Papers with diagnostic tasks

7 papers
Benchmarking and Mitigating the Impact of Noisy User Prompts in Medical VLMs via Cross-Modal Reflection (2026.eacl-industry)

Copied to clipboard

Challenge: Existing medical vision-language models follow user-provided prompts blindly, a new study finds . current models are noisy, causing problems with reliability in real-world interactions .
Approach: They propose a method to evaluate the influence of clinical prompts on medical vision-language models . they use cross-modal reflection chain-of-thought to train the model to produce reasoning paths .
Outcome: The proposed method significantly improves the robustness against noisy prompts . existing Med-VLMs follow user-provided prompts blindly, the authors show .
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale (2026.findings-eacl)

Copied to clipboard

Challenge: Existing approaches to evaluate language models fail to provide structural clarity and verifiable inference.
Approach: They propose to use a large-scale dataset of programmatically verified reasoning traces to evaluate structured logical inference.
Outcome: The proposed model achieves 45.7% accuracy on masked operation prediction and 27% on two-step completion.
Improving Compositional Generalization with Latent Structure and Data Augmentation (2022.naacl-main)

Copied to clipboard

Challenge: Generic unstructured neural networks struggle on out-of-distribution compositional generalization.
Approach: They propose a method to recombinate examples from a model called Compositional Structure Learner and add them to a pre-trained sequence-to-sequence model.
Outcome: The proposed model is even stronger than a T5-CSL ensemble on two real world compositional generalization tasks.
Good-Enough Compositional Data Augmentation (2020.acl-main)

Copied to clipboard

Challenge: a proposed data augmentation protocol provides a compositional inductive bias in conditional and unconditional sequence models.
Approach: They propose a data augmentation protocol that provides a compositional inductive bias in conditional and unconditional sequence models by replacing discontinuous fragments with other fragments that appear in at least one similar environment.
Outcome: The proposed protocol reduces error rate by 87% on diagnostic tasks and 16% on semantic parsing tasks.
Beyond Static Profiles: Capturing the Fluidity of User Preferences in Diverse Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to personalize Large Language Models often default to homogeneous behaviors . preferences can shift, and conflict, depending on context, authors argue .
Approach: They propose a hierarchical taxonomy to differentiate between stable and situational preferences . they use a dataset of 10k meticulously curated preferences to test their taxonomies .
Outcome: The proposed model differentiates between stable and situational preferences based on curated user preferences . it provides a practical testbed for advancing dynamic, context-aware personalization in conversational agents.
Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational Knowledge (2025.emnlp-main)

Copied to clipboard

Challenge: Using large language models, we study their morphosyntactic competence and generalization capabilities.
Approach: They propose to use morphosyntactic tasks to study their linguistic knowledge and generalization capabilities to extract different types of morphological structure for typologically diverse languages.
Outcome: The proposed models outperform GPT-4o and LLaMA 3.3-70B in all diagnostic tasks, but show little evidence of abstract morphological rule learning.
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form–Meaning Mapping (2026.acl-long)

Copied to clipboard

Challenge: a visual Iconicity test is used to evaluate vision–language models based on visual form and iconicity ratings.
Approach: They propose a video-based benchmark to evaluate vision–language models on three tasks . they assess 17 state-of-the-art VLMs in zero- and few-shot settings on Sign Language of the Netherlands .
Outcome: The proposed benchmark evaluates 17 state-of-the-art VLMs on Sign Language of the Netherlands . they achieve moderate to strong alignment with human iconicity ratings, but fail to infer lexical meaning from visual form alone .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations